Papers with text-only approaches
A Character-Centric Creative Story Generation via Imagination (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing narrative generation models lack diversity and character depth, but they are inadequate for human creativity. |
| Approach: | They propose a novel story generation framework called CCI that leverages images to create stories that are diverse and creative in their themes and richer in content. |
| Outcome: | The proposed framework significantly improves various aspects of the stories’ creativity. |
SpeechT-RAG: Reliable Depression Detection in LLMs with Retrieval-Augmented Generation Using Speech Timing Information (2025.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have been extensively utilized for health-related tasks, yet their performance in depression detection remains limited when relying solely on text input. |
| Approach: | They propose a system that leverages speech timing features for depression detection and reliable confidence estimation. |
| Outcome: | The proposed system outperforms text-based RAG systems in depression detection and confidence estimation. |
What if Othello-Playing Language Models Could See? (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a multi-modal model trained on move sequences and board images is a popular testbed for language models . |
| Approach: | They propose a multi-modal model trained jointly on move sequences and board images. |
| Outcome: | The proposed multi-modal model trains on move sequences and board images. |
MAVL: A Multilingual Audio-Video Lyrics Dataset for Animated Song Translation (2025.emnlp-main)
Copied to clipboard
| Challenge: | Experimental results show that multimodal, multimodal approaches to lyrics translation are more effective than text-only approaches. |
| Approach: | They propose a multilingual, multimodal benchmark for singable lyrics translation . they propose syllable-constrained audio-video LLM with Chain-of-Thought . |
| Outcome: | The proposed system outperforms text-based models in singability and contextual accuracy. |